Papers with decision accuracy

9 papers
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain (2024.findings-acl)

Copied to clipboard

Challenge: a new multimodal decision-making benchmark evaluates the integrated capabilities of multimodal large language models.
Approach: They propose a multimodal decision-making benchmark for evaluating MLLMs . they propose an automatic evaluation protocol to assess 10 prevalent ML models .
Outcome: The proposed benchmark improves performance of multimodal large language models in three scenarios . the model is required to integrate multiple capabilities to make accurate decisions .
DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) tuned for safety often avoid acknowledging demographic differences . current safety alignment forces LLMs to default to identity-blindness even when demographic differences are factually correct or contextually justified.
Approach: They propose a tool to classify whether a correct answer requires recognizing group differences . they use label-conditioned reasoning from a teacher to audit outputs for harm drift cases .
Outcome: The proposed model improves accuracy and safety on eight benchmarks.
KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in GUI agents have limited app-specific knowledge of complex mobile tasks.
Approach: They propose a Knowledge Graph-driven Retrieval-Augmented Generation framework that transforms fragmented UTGs into structured vector databases for efficient real-time retrieval.
Outcome: The proposed framework outperforms existing methods in a 75.8% success rate and 84.6% decision accuracy test across mobile apps.
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies on the impact of feedback on human decision-making are limited as people are not equipped to assess the quality of AI predictions.
Approach: They compare the quality of MT inputs and outputs with explicit and implicit feedbacks that directly give users an assessment of translation quality using error highlights and LLM explanations.
Outcome: The proposed model improves decision accuracy and appropriate reliance by using error highlights and explanations, and by using backtranslation and question–answer tables.
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing vision-language planning methods struggle with long-horizon reasoning in dynamic environments due to the difficulty of training models to generate high-quality reasoning processes.
Approach: They propose a framework that enhances reasoning and action selection for long-horizon task planning through structured evaluation and optimized training.
Outcome: The proposed framework outperforms existing methods on short-horizon tasks but struggles with long-horizon reasoning in dynamic environments.
AskQE: Question Answering as Automatic Evaluation for Machine Translation (2025.findings-acl)

Copied to clipboard

Challenge: Existing MT error detection and quality estimation (QE) techniques do not address this practical scenario.
Approach: They propose a question generation and answering framework that detects critical MT errors and provides actionable feedback to help users decide whether to accept or reject MT outputs even without the knowledge of the target language.
Outcome: The proposed framework has higher Kendall’s Tau correlation and decision accuracy with human ratings compared to other QE metrics.
InsLogicBench: An Argumentation Logic Grounded Benchmark for Complex Insurance Claims Adjudication (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for insurance claims adjudication are limited to information retrieval or simple multiple-choice setups.
Approach: They propose a benchmark that provides complete reasoning traces linking factual inputs, relevant policy clauses, and final verdicts.
Outcome: The proposed benchmark shows that models often produce correct decisions but fail to provide precise justifications, highlighting a critical discrepancy between decision accuracy and logical reasoning capabilities.
Decisive: Guiding User Decisions with Optimal Preference Elicitation from Unstructured Documents (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for analyzing information from multiple sources are often too complex or fail to capture nuanced preferences accurately.
Approach: They propose an interactive decision-making framework that combines document-grounded reasoning with Bayesian preference inference.
Outcome: The proposed approach outperforms general-purpose LLMs and existing decision-support systems in achieving up to 20% improvement in decision accuracy over strong baselines across domains.
DAC-Bench: A Decision-Aware Benchmark for Compositional Mobile GUI Tasks (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on short, linear workflows and step-level accuracy, highlighting performance degradations.
Approach: They propose a decision-aware benchmark with compositional tasks comprising 830 episodes and 11,345 action steps across 35 applications on Android and iOS.
Outcome: The proposed benchmarks show performance degradation and branch correctness issues in 7 different GUI agents.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations